Interoperability Layer
The interoperability layer (IOL) is the single door into a health information exchange. Every point-of-service system talks to it; none talks directly to a registry or to another point-of-service system.
It exists to make integration cost linear instead of quadratic, and to put authentication, audit and error handling in one place that can be operated properly rather than in twenty places that cannot.
What it does
point-of-service systems
│
▼
┌──────────────────────────────────────────────┐
│ 1. Authenticate the caller │
│ 2. Authorise the transaction │
│ 3. Validate the payload │
│ 4. Resolve identity (client, facility, HW) │
│ 5. Transform format and terminology │
│ 6. Orchestrate multi-step transactions │
│ 7. Route to the target service(s) │
│ 8. Persist an audit record │
│ 9. Return a meaningful response or error │
└──────────────────────────────────────────────┘
│
▼
registries · SHR · HMIS · domain services
Authentication
Every caller has a distinct machine identity — a client certificate, an OAuth client credential, or an API key issued per system and per environment. Shared credentials destroy the audit trail, which is usually the reason the layer was funded.
Interlinking and entity mapping
Incoming data carries local identifiers. Before it is stored or forwarded, the IOL resolves them against the client registry, facility registry and health worker registry, and either substitutes or adds the shared identifier.
This is the most valuable thing the layer does. It is also where transactions legitimately fail — an unresolvable patient is a business exception, not a technical error, and needs a human queue rather than a retry loop.
Transformation
Format translation (HL7 v2 ⇄ FHIR ⇄ CSV ⇄ custom XML) and terminology translation (local codes → LOINC, SNOMED CT, ICD) using a terminology service rather than hard-coded maps.
Hard-coded maps are the single most common source of decay in these
architectures. A ConceptMap in a terminology service can be updated by a
terminologist; a switch statement in a mediator can only be updated by
redeploying software.
Orchestration
Some transactions need several steps: resolve the patient, fetch prior results, save the encounter, notify the surveillance service. The IOL sequences them and decides what happens if step three fails after steps one and two succeeded.
Keep orchestration explicit and idempotent. Every operation should be safe to retry, which in practice means client-supplied transaction identifiers and upsert semantics rather than blind inserts.
Mediation
A mediator is a small, purpose-built unit of logic for one integration — one system, one use case. Mediators keep the layer's core generic and let integrations be developed, versioned and retired independently.
Audit
An immutable record of every transaction: who called, on whose behalf, for which
patient, for what purpose, what was returned, and when. This is what a data
protection authority asks for, and what makes a breach investigable. See
consent and trust and FHIR
AuditEvent.
Audit records contain personal data. They need their own retention policy, access control and, usually, their own storage.
Monitoring
Message volume, error rates by channel, latency percentiles, queue depth, and per-system availability. See observability. The operational question this answers is which system is broken right now, which the layer is uniquely positioned to know.
What it must not do
An interoperability layer that acquires these responsibilities becomes unmaintainable:
- Store the clinical record. That is the shared health record.
- Hold business rules. Clinical decision logic belongs in a decision support service, not in a routing mediator.
- Become the system of record for identity. It resolves identity; the registry owns it.
- Silently repair bad data. Normalising a date format is fine; inventing a missing facility code is not — it hides a problem that will surface as wrong statistics later.
Failure modes
| Failure | Symptom | Mitigation |
|---|---|---|
| Single point of failure | All exchange stops when the IOL is down | Redundant instances, health checks, store-and-forward at the edge |
| Synchronous coupling | One slow registry slows every transaction | Timeouts, circuit breakers, asynchronous patterns for non-urgent flows |
| Queue backlog | Silent data loss or hours-old data | Alert on queue depth and age, not just on errors |
| Mediator sprawl | Dozens of unversioned mediators nobody owns | Ownership register, deprecation policy, conformance tests |
| Credential sharing | Audit trail cannot attribute actions | One identity per system per environment, rotated |
| Hidden transformation logic | Data changes and nobody knows where | Terminology maps in a terminology service; transformations under version control |
| No error contract | Failures accumulate in a log | Defined error codes, a triage queue and a named human owner |
Implementations
| Option | Notes | Tier |
|---|---|---|
| OpenHIM | The OpenHIE reference implementation. Purpose-built for this role: channel routing, mediator framework, transaction log, console. | 2 |
| Mirth / NextGen Connect | Mature health integration engine, strong HL7 v2 handling. | 2 |
| Apache Camel | General-purpose integration framework with health components; maximum flexibility, more to build. | 2 |
| Apache NiFi | Flow-based data routing; better suited to bulk and analytics pipelines than transactional exchange. | 2 |
| OpenFn | Workflow automation and integration, used in several global-health deployments. | 2 |
| API gateway (Kong, Traefik, Envoy) + services | Handles authentication, routing and rate limiting well; transformation and orchestration must be built. | 2 |
A gateway is not an interoperability layer on its own — it does steps 1, 2, 7 and part of 8 above, and none of 4, 5 or 6. Combining a gateway with dedicated mediation services is a legitimate architecture; assuming the gateway is sufficient is not.
See integration engines for a fuller comparison.
Sizing and operations
Before choosing, answer:
- Peak throughput — laboratory results and immunisation campaigns produce bursts, not averages
- Latency budget — a patient-identity lookup during registration has a very different budget from an overnight HMIS submission
- Availability target — and what point-of-service systems do when it is not met
- Who operates it — the layer needs 24/7 ownership; if no team exists, the architecture is aspirational
References
- OpenHIE architecture — https://ohie.org/
- OpenHIM — https://openhim.org/
- HL7 FHIR
AuditEvent— https://hl7.org/fhir/auditevent.html - IHE ATNA (audit trail and node authentication) — https://www.ihe.net/